Papers with LLM factuality

3 papers
LEAF: Learning and Evaluation Augmented by Fact-Checking to Improve Factualness in Large Language Models (2025.emnlp-industry)

Copied to clipboard

Challenge: Large language models (LLMs) struggle with factual accuracy in knowledge-intensive domains like healthcare.
Approach: They propose a framework for improving LLM factuality in medical question answering . RAFE, Fact-Check-then-RAG and Learning from Fact Check are components .
Outcome: Experimental results show that LEAF outperforms Factcheck-GPT in detecting inaccuracies and corrects errors without labeling . the framework provides a scalable solution for industrial applications requiring high factuality scores.
When Benchmarks Age: Temporal Misalignment through Large Language Model Factuality Evaluation (2026.eacl-short)

Copied to clipboard

Challenge: Existing studies on LLM factuality evaluation have not investigated the reliability of static evaluation benchmarks.
Approach: They examine five popular factuality benchmarks and eight LLMs released over different years to assess their reliability.
Outcome: The proposed method compared five popular factuality benchmarks and eight LLMs released over different years.
How Does Response Length Affect Long-Form Factuality (2025.findings-acl)

Copied to clipboard

Challenge: Despite growing attention to LLM factuality, the effect of response length on factual accuracy remains underexplored.
Approach: They propose an automatic and bi-level long-form factuality evaluation framework which achieves high agreement with human annotations while being cost-effective.
Outcome: The proposed framework achieves high agreement with human annotations while being cost-effective.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations